Papers with quality assurance

13 papers
Synthetic Data for Evaluation: Supporting LLM-as-a-Judge Workflows with EvalAssist (2025.emnlp-demos)

Copied to clipboard

Challenge: EvalAssist is a web-based application designed to assist human-centered evaluation of language model outputs.
Approach: They propose a synthetic data generation tool integrated into EvalAssist to assist human-centered evaluation of language model outputs.
Outcome: The proposed tool supports flexible prompting, RAG-based grounding, persona diversity, and iterative generation workflows.
PromptLab: A Collaborative Platform for Prompt Engineering and Dataset Curation (2026.eacl-demo)

Copied to clipboard

Challenge: PromptLab is a web-based prompt engineering platform for collaborative prompt development across diverse natural language processing tasks and datasets.
Approach: They propose to integrate prompt generation via OpenRouter and provide real-time validation with multiple Large Language Models.
Outcome: The platform addresses primary challenges in prompt development, including template creation, collaborative review, and quality assurance through a comprehensive workflow that supports both individual researchers and team-based projects.
Sensing and Learning Human Annotators Engaged in Narrative Sensemaking (N18-4)

Copied to clipboard

Challenge: a substantial sector of the gig economy is the use of crowdworkers to annotate data for machine learning and analysis.
Approach: They propose a narrative-sorting annotation task that sorts tweets chronologically by topic, emotional content, and length.
Outcome: The proposed task enables readers to sort sequential, target-topical, and emotionally emotional tweets.
AI Coach Assist: An Automated Approach for Call Recommendation in Contact Centers for Agent Coaching (2023.acl-industry)

Copied to clipboard

Challenge: In recent years, the utilization of Artificial Intelligence (AI) in the contact center industry is on the rise.
Approach: They present a transformer-based pairwise sentence classification model that analyzes call transcripts to determine which calls are most relevant for coaching purposes.
Outcome: The proposed model can determine which calls are most relevant for coaching purposes based on quality assurance queries/questions asked by managers or supervisors .
Lost and Found: Computational Quality Assurance of Crowdsourced Knowledge on Morphological Defectivity in Wiktionary (2025.acl-srw)

Copied to clipboard

Challenge: a recent study shows that wikis are not reliable for linguistic knowledge of defects in understudied languages.
Approach: They customize a neural morphological analyzer to annotate Latin and Italian corpora . they validated morphology using crowd-sourced data from Wiktionary to find defects .
Outcome: The proposed algorithm annotates Latin and Italian corpora using crowd-sourced data . results show that 7% of Latin lemmata listed as defective show strong corpus evidence of being non-defective.
Translation Crowdsourcing: Creating a Multilingual Corpus of Online Educational Content (L18-1)

Copied to clipboard

Challenge: a large corpus of online content has been developed via large-scale crowdsourcing.
Approach: They describe a multilingual corpus of online content that has been manually translated into 11 European and BRIC languages using the crowdsourcing platform.
Outcome: The proposed corpus is a product of the EU-funded TraMOOC project and is used to train, tune and test machine translation engines.
Experience Report: Implementing Machine Translation in a Regulated Industry (2025.emnlp-industry)

Copied to clipboard

Challenge: a global medical technology company has invested substantial resources in translating content into the various languages required across their global markets.
Approach: They propose to use human-in-the-loop validation to evaluate machine translation systems in a medical technology company.
Outcome: The proposed method dominates reviewer preference across all languages and tones of interest, the authors show . the "Gold" control ranks poorly in one language and the lower ranks have high variance.
Self-prompted Chain-of-Thought on Large Language Models for Open-domain Multi-hop Reasoning (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing open-domain question-answering methods lack quality assurance . existing methods lack scalability and poor diversity, hindering LLMs' capabilities .
Approach: They propose an open-domain multi-hop reasoning framework to answer multi-choice questions . they propose an adaptive sampler for in-context selection and self-prompted inference .
Outcome: The proposed framework surpasses the existing SOTA methods on large-scale datasets and doubles the zero-shot performance of small-scale LLMs.
BertNet: Harvesting Knowledge Graphs with Arbitrary Relations from Pretrained Language Models (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to construct knowledge graphs are limited to a small set of relations due to manual cost or restrictions in text corpus.
Approach: They propose to automatically construct knowledge graphs (KGs) of diverse new relations from pretrained language models that accept knowledge queries with prompts.
Outcome: The proposed framework extracts knowledge of over 400 new relations from pretrained language models, including RoBERTaNet, with minimal input of a relation definition and a few shot of example entity pairs.
FRAME: Feedback-Refined Agent Methodology for Enhancing Medical Research Insights (2025.findings-acl)

Copied to clipboard

Challenge: Existing approaches to automate scientific research are limited by human cognitive constraints and timeintensive workflows.
Approach: They propose a framework that enhances medical paper generation through iterative refinement and structured feedback.
Outcome: The proposed framework achieves significant improvements over conventional methods across multiple models and evaluation dimensions.
Intrinsic Evaluation of Summarization Datasets (2020.emnlp-main)

Copied to clipboard

Challenge: Almost all popular summarization datasets do not come with inherent quality assurance guarantees.
Approach: They propose to use 5 metrics to evaluate quality of summarization datasets . they find that data usage in recent summarizing research is inconsistent with the properties of the data.
Outcome: The proposed metrics can be inexpensive heuristics for detecting generically low quality examples.
PMIndiaSum: Multilingual and Cross-lingual Headline Summarization for Languages in India (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing datasets for Indian languages are limited in terms of coverage and size.
Approach: They propose a multilingual and massively parallel summarization corpus focused on languages in India that provides a training and testing ground for four language families, 14 languages, and the largest to date with 196 language pairs.
Outcome: The proposed dataset provides a training and testing ground for four language families, 14 languages, and the largest to date with 196 language pairs.
LongMP-Bench: A Benchmark for Multimodal Persona Understanding in Long-Term Dialogues (2026.findings-acl)

Copied to clipboard

Challenge: Existing datasets suffer from limited persona diversity and static, overly simplified settings, making them insufficient for capturing the complexity of real-world interactions.
Approach: They propose a benchmark to evaluate models' ability to understand evolving user personas within long-term multimodal dialogues by using a dataset that contains long conversations from 150 users.
Outcome: The proposed benchmark aims to assess models' ability to track persona evolution, integrate visual and textual inputs, and apply persona understanding in realistic dialogue scenarios.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations